Papers with verbatim memorization
Copyright Violations and Large Language Models (2023.emnlp-main)
Copied to clipboard
| Challenge: | a recent study examines the extent to which language models can memorize training data . a fair use exemption to copyright laws allows for limited use of copyrighted material . |
| Approach: | They examine the extent to which language models can redistribute copyrighted text . they use a range of popular books and coding problems to study copyright violations . |
| Outcome: | This study examines the extent to which language models can redistribute copyrighted text . it shows that language models may memorize entire chunks of training data . |
Demystifying Verbatim Memorization in Large Language Models (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies have shown that Large Language Models (LLMs) memorize long sequences verbatim, with serious copyright and privacy implications. |
| Approach: | They develop a framework to study verbatim memorization in a controlled setting by continuing pre-training from Pythia checkpoints with injected sequences. |
| Outcome: | The proposed framework creates a control model M () and a treatment model M with injected sequences. |